In situ data processing for extreme scale computing
نویسندگان
چکیده
The DOE leadership facilities were established in 2004 to provide scientist capability computing for high-profile science. Since it’s inception, the systems went from 14 TF to 1.8 PF, an increase of 100 in 5 years, and will increase by another factor of 100 in 5 more years. This growth, along with user policies, which enable scientist to run at, scale for long periods of time, have allowed scientist to write unprecedented amounts of data to the file system. In the same time, the effective speed of the I/O system (time to write full system memory to the file system) has decreased, going from 49 TB/140 GB/s on ASCI purple, to 300 TB/200 GB/s on jaguar. As future systems will further intensify this imbalance, we need to extend the definition of I/O to include that of I/O pipelines, which blend together; 1) Self describing file formats, 2) Staging methods with visualization and analysis “plug-ins”, 3) new techniques to multiplex outputs using a novel Service Oriented Approach (SOA), in a 4) Easy-to-use I/O componentization framework which can add resiliency to large scale calculations. Our approach to this has been in the community created ADIOS system, driven by many of the leading edge application teams, and by many of the top I/O researchers.
منابع مشابه
Outlier Detection Using Extreme Learning Machines Based on Quantum Fuzzy C-Means
One of the most important concerns of a data miner is always to have accurate and error-free data. Data that does not contain human errors and whose records are full and contain correct data. In this paper, a new learning model based on an extreme learning machine neural network is proposed for outlier detection. The function of neural networks depends on various parameters such as the structur...
متن کاملIn Situ Methods, Infrastructures, and Applications on High Performance Computing Platforms
The considerable interest in the high performance computing (HPC) community regarding analyzing and visualization data without first writing to disk, i.e., in situ processing, is due to several factors. First is an I/O cost savings, where data is analyzed/visualized while being generated, without first storing to a filesystem. Second is the potential for increased accuracy, where fine temporal ...
متن کاملEntropy-based Consensus for Distributed Data Clustering
The increasingly larger scale of available data and the more restrictive concerns on their privacy are some of the challenging aspects of data mining today. In this paper, Entropy-based Consensus on Cluster Centers (EC3) is introduced for clustering in distributed systems with a consideration for confidentiality of data; i.e. it is the negotiations among local cluster centers that are used in t...
متن کاملCloud Computing Technology Algorithms Capabilities in Managing and Processing Big Data in Business Organizations: MapReduce, Hadoop, Parallel Programming
The objective of this study is to verify the importance of the capabilities of cloud computing services in managing and analyzing big data in business organizations because the rapid development in the use of information technology in general and network technology in particular, has led to the trend of many organizations to make their applications available for use via electronic platforms hos...
متن کاملNumerical modelling of the underground roadways in coal mines– uncertainties caused by use of empirical-based downgrading methods and in situ stresses
Numerical modelling techniques are not new for mining industry and civil engineering projects anymore. These techniques have been widely used for rock engineering problems such as stability analysis and support design of roadways and tunnels, caving and subsidence prediction, and stability analysis of rock slopes. Despite the significant advancement in the computational mechanics and availabili...
متن کاملA characterization of workflow management systems for extreme-scale applications
Automation of the execution of computational tasks is at the heart of improving scientific productivity. Over the last years, scientific workflows have been established as an important abstraction that captures data processing and computation of large and complex scientific applications. By allowing scientists to model and express entire data processing steps and their dependencies, workflow ma...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2011